Micron Document
`:top
In `F33f`_`[probability`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Probability_theory]`_`f, and `F33f`_`[statistics`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Statistics]`_`f, a `!multivariate random variable`! or `!random vector`! is a list or `F33f`_`[vector`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Vector_(mathematics)]`_`f of mathematical `F33f`_`[variables`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Variable_(mathematics)]`_`f each of whose value is unknown, either because the value has not yet occurred or because there is imperfect knowledge of its value. The individual variables in a random vector are grouped together because they are all part of a single mathematical system — often they represent different properties of an individual `F33f`_`[statistical unit`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Statistical_unit]`_`f. For example, while a given person has a specific age, height and weight, the representation of these features of `*an unspecified person`* from within a group would be a random vector. Normally each element of a random vector is a `F33f`_`[real number`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Real_number]`_`f.

Random vectors are often used as the underlying implementation of various types of aggregate `F33f`_`[random variables`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Random_variable]`_`f, e.g. a `F33f`_`[random matrix`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Random_matrix]`_`f, `F33f`_`[random tree`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Random_tree]`_`f, `F33f`_`[random sequence`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Random_sequence]`_`f, `F33f`_`[stochastic process`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Stochastic_process]`_`f, etc.

Formally, a multivariate random variable is a `F33f`_`[column vector`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Column_vector]`_`f X = ( X 1 , … … , X n ) T {\\displaystyle \\mathbf {X} =(X_{1},\\dots ,X_{n})^{\\mathsf {T}}} (or its `F33f`_`[transpose`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Transpose]`_`f, which is a `F33f`_`[row vector`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Row_vector]`_`f) whose components are `F33f`_`[random variables`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Random_variable]`_`f on the `F33f`_`[probability space`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Probability_space]`_`f ( Ω Ω , F , P ) {\\displaystyle (\\Omega ,{\\mathcal {F}},P)} , where Ω Ω {\\displaystyle \\Omega } is the `F33f`_`[sample space`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Sample_space]`_`f, F {\\displaystyle {\\mathcal {F}}} is the `F33f`_`[sigma-algebra`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Sigma-algebra]`_`f (the collection of all events), and P {\\displaystyle P} is the `F33f`_`[probability measure`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Probability_measure]`_`f (a function returning each event's `F33f`_`[probability`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Probability]`_`f).

>>Contents

• `F0af`_`[Probability distribution`#probability-distribution]`_`f
• `F0af`_`[Operations on random vectors`#operations-on-random-vectors]`_`f
• `F0af`_`[Affine transformations`#affine-transformations]`_`f
• `F0af`_`[Invertible mappings`#invertible-mappings]`_`f
• `F0af`_`[Expected value`#expected-value]`_`f
• `F0af`_`[Covariance and cross-covariance`#covariance-and-cross-covariance]`_`f
• `F0af`_`[Definitions`#definitions]`_`f
• `F0af`_`[Properties`#properties]`_`f
• `F0af`_`[Uncorrelatedness`#uncorrelatedness]`_`f
• `F0af`_`[Correlation and cross-correlation`#correlation-and-cross-correlation]`_`f
• `F0af`_`[Definitions`#definitions]`_`f
• `F0af`_`[Properties`#properties]`_`f
• `F0af`_`[Orthogonality`#orthogonality]`_`f
• `F0af`_`[Independence`#independence]`_`f
• `F0af`_`[Characteristic function`#characteristic-function]`_`f
• `F0af`_`[Further properties`#further-properties]`_`f
• `F0af`_`[Expectation of a quadratic form`#expectation-of-a-quadratic-form]`_`f
• `F0af`_`[Expectation of the product of two different quadratic forms`#expectation-of-the-product-of-two-different-quadratic-forms]`_`f
• `F0af`_`[Applications`#applications]`_`f
• `F0af`_`[Portfolio theory`#portfolio-theory]`_`f
• `F0af`_`[Regression theory`#regression-theory]`_`f
• `F0af`_`[Vector time series`#vector-time-series]`_`f
• `F0af`_`[References`#references]`_`f
• `F0af`_`[Further reading`#further-reading]`_`f

-─

>>Probability distribution

Every random vector gives rise to a probability measure on R n {\\displaystyle \\mathbb {R} ^{n}} with the `F33f`_`[Borel algebra`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Borel_algebra]`_`f as the underlying sigma-algebra. This measure is also known as the `F33f`_`[joint probability distribution`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Joint_probability_distribution]`_`f, the joint distribution, or the multivariate distribution of the random vector.

The `F33f`_`[distributions`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Probability_distribution]`_`f of each of the component random variables X i {\\displaystyle X_{i}} are called `F33f`_`[marginal distributions`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Marginal_distribution]`_`f. The `F33f`_`[conditional probability distribution`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Conditional_probability_distribution]`_`f of X i {\\displaystyle X_{i}} given X j {\\displaystyle X_{j}} is the probability distribution of X i {\\displaystyle X_{i}} when X j {\\displaystyle X_{j}} is known to be a particular value.

The `!cumulative distribution function`! F X : R n ↦ ↦ [ 0 , 1 ] {\\displaystyle F_{\\mathbf {X} }:\\mathbb {R} ^{n}\\mapsto [0,1]} of a random vector X = ( X 1 , … … , X n ) T {\\displaystyle \\mathbf {X} =(X_{1},\\dots ,X_{n})^{\\mathsf {T}}} is defined as`:cite-ref-gallager-1-0[`F5bf`_`[1`#cite-note-gallager-1]`_`f]

where x = ( x 1 , … … , x n ) T {\\displaystyle \\mathbf {x} =(x_{1},\\dots ,x_{n})^{\\mathsf {T}}} .

>>Operations on random vectors

Random vectors can be subjected to the same kinds of `F33f`_`[algebraic operations`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Euclidean_vector]`_`f as can non-random vectors: addition, subtraction, multiplication by a `F33f`_`[scalar`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Scalar_(mathematics)]`_`f, and the taking of `F33f`_`[inner products`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Dot_product]`_`f.

>>>Affine transformations

Similarly, a new random vector Y {\\displaystyle \\mathbf {Y} } can be defined by applying an `F33f`_`[affine transformation`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Affine_transformation]`_`f g : : R n → → R n {\\displaystyle g\\colon \\mathbb {R} ^{n}\\to \\mathbb {R} ^{n}} to a random vector X {\\displaystyle \\mathbf {X} } :

Y = A X + b {\\displaystyle \\mathbf {Y} =\\mathbf {A} \\mathbf {X} +b} , where A {\\displaystyle \\mathbf {A} } is an n × × n {\\displaystyle n\\times n} matrix and b {\\displaystyle b} is an n × × 1 {\\displaystyle n\\times 1} column vector.

If A {\\displaystyle \\mathbf {A} } is an `F33f`_`[invertible matrix`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Invertible_matrix]`_`f and X {\\displaystyle \\textstyle \\mathbf {X} } has a probability density function f X {\\displaystyle f_{\\mathbf {X} }} , then the probability density of Y {\\displaystyle \\mathbf {Y} } is

f Y ( y ) = f X ( A − − 1 ( y − − b ) ) | det A | {\\displaystyle f_{\\mathbf {Y} }(y)={\\frac {f_{\\mathbf {X} }(\\mathbf {A} ^{-1}(y-b))}{|\\det \\mathbf {A} |}}} .

>>>Invertible mappings

More generally we can study invertible mappings of random vectors.`:cite-ref-lapidoth-2-0[`F5bf`_`[2`#cite-note-lapidoth-2]`_`f]

Let g {\\displaystyle g} be a one-to-one mapping from an open subset D {\\displaystyle {\\mathcal {D}}} of R n {\\displaystyle \\mathbb {R} ^{n}} onto a subset R {\\displaystyle {\\mathcal {R}}} of R n {\\displaystyle \\mathbb {R} ^{n}} , let g {\\displaystyle g} have continuous partial derivatives in D {\\displaystyle {\\mathcal {D}}} and let the `F33f`_`[Jacobian determinant`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Jacobian_matrix_and_determinant]`_`f det ( ∂ ∂ y ∂ ∂ x ) {\\displaystyle \\det \\left({\\frac {\\partial \\mathbf {y} }{\\partial \\mathbf {x} }}\\right)} of g {\\displaystyle g} be zero at no point of D {\\displaystyle {\\mathcal {D}}} . Assume that the real random vector X {\\displaystyle \\mathbf {X} } has a probability density function f X ( x ) {\\displaystyle f_{\\mathbf {X} }(\\mathbf {x} )} and satisfies P ( X ∈ ∈ D ) = 1 {\\displaystyle P(\\mathbf {X} \\in {\\mathcal {D}})=1} . Then the random vector Y = g ( X ) {\\displaystyle \\mathbf {Y} =g(\\mathbf {X} )} is of probability density

f Y ( y ) = f X ( x ) | det ( ∂ ∂ y ∂ ∂ x ) | | x = g − − 1 ( y ) 1 ( y ∈ ∈ R Y ) {\\displaystyle \\left.f_{\\mathbf {Y} }(\\mathbf {y} )={\\frac {f_{\\mathbf {X} }(\\mathbf {x} )}{\\left|\\det \\left({\\frac {\\partial \\mathbf {y} }{\\partial \\mathbf {x} }}\\right)\\right|}}\\right|_{\\mathbf {x} =g^{-1}(\\mathbf {y} )}\\mathbf {1} (\\mathbf {y} \\in R_{\\mathbf {Y} })}

where 1 {\\displaystyle \\mathbf {1} } denotes the `F33f`_`[indicator function`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Indicator_function]`_`f and set R Y = { y = g ( x ) : f X ( x ) > 0 } ⊆ ⊆ R {\\displaystyle R_{\\mathbf {Y} }=\\{\\mathbf {y} =g(\\mathbf {x} ):f_{\\mathbf {X} }(\\mathbf {x} )>0\\}\\subseteq {\\mathcal {R}}} denotes support of Y {\\displaystyle \\mathbf {Y} } .

>>Expected value

The `F33f`_`[expected value`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Expected_value]`_`f or mean of a random vector X {\\displaystyle \\mathbf {X} } is a fixed vector E ⁡ ⁡ [ X ] {\\displaystyle \\operatorname {E} [\\mathbf {X} ]} whose elements are the expected values of the respective random variables.`:cite-ref-gubner-3-0[`F5bf`_`[3`#cite-note-gubner-3]`_`f]

>>Covariance and cross-covariance

>>>Definitions

The `!`F33f`_`[covariance matrix`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Covariance_matrix]`_`f`! (also called `!second central moment`! or variance-covariance matrix) of an n × × 1 {\\displaystyle n\\times 1} random vector is an n × × n {\\displaystyle n\\times n} `F33f`_`[matrix`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Matrix_(mathematics)]`_`f whose (`*i,j`*)th element is the `F33f`_`[covariance`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Covariance]`_`f between the `*i`* th and the `*j`* th random variables. The covariance matrix is the expected value, element by element, of the n × × n {\\displaystyle n\\times n} matrix `F33f`_`[computed as`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Matrix_multiplication]`_`f [ X − − E ⁡ ⁡ [ X ] ] [ X − − E ⁡ ⁡ [ X ] ] T {\\displaystyle [\\mathbf {X} -\\operatorname {E} [\\mathbf {X} ]][\\mathbf {X} -\\operatorname {E} [\\mathbf {X} ]]^{T}} , where the superscript T refers to the transpose of the indicated vector:`:cite-ref-lapidoth-2-1[`F5bf`_`[2`#cite-note-lapidoth-2]`_`f]`:cite-ref-gubner-3-1[`F5bf`_`[3`#cite-note-gubner-3]`_`f]

By extension, the `!`F33f`_`[cross-covariance matrix`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Cross-covariance_matrix]`_`f`! between two random vectors X {\\displaystyle \\mathbf {X} } and Y {\\displaystyle \\mathbf {Y} } ( X {\\displaystyle \\mathbf {X} } having n {\\displaystyle n} elements and Y {\\displaystyle \\mathbf {Y} } having p {\\displaystyle p} elements) is the n × × p {\\displaystyle n\\times p} matrix`:cite-ref-gubner-3-2[`F5bf`_`[3`#cite-note-gubner-3]`_`f]

where again the matrix expectation is taken element-by-element in the matrix. Here the (`*i,j`*)th element is the covariance between the `*i`* th element of X {\\displaystyle \\mathbf {X} } and the `*j`* th element of Y {\\displaystyle \\mathbf {Y} } .

>>>Properties

The covariance matrix is a `F33f`_`[symmetric matrix`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Symmetric_matrix]`_`f, i.e.`:cite-ref-lapidoth-2-2[`F5bf`_`[2`#cite-note-lapidoth-2]`_`f]

K X X T = K X X {\\displaystyle \\operatorname {K} _{\\mathbf {X} \\mathbf {X} }^{T}=\\operatorname {K} _{\\mathbf {X} \\mathbf {X} }} .

The covariance matrix is a `F33f`_`[positive semidefinite matrix`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Positive_semidefinite_matrix]`_`f, i.e.`:cite-ref-lapidoth-2-3[`F5bf`_`[2`#cite-note-lapidoth-2]`_`f]

a T K X X ⁡ ⁡ a ≥ ≥ 0 for all a ∈ ∈ R n {\\displaystyle \\mathbf {a} ^{T}\\operatorname {K} _{\\mathbf {X} \\mathbf {X} }\\mathbf {a} \\geq 0\\quad {\\text{for all }}\\mathbf {a} \\in \\mathbb {R} ^{n}} .

The cross-covariance matrix Cov ⁡ ⁡ [ Y , X ] {\\displaystyle \\operatorname {Cov} [\\mathbf {Y} ,\\mathbf {X} ]} is simply the transpose of the matrix Cov ⁡ ⁡ [ X , Y ] {\\displaystyle \\operatorname {Cov} [\\mathbf {X} ,\\mathbf {Y} ]} , i.e.

K Y X = K X Y T {\\displaystyle \\operatorname {K} _{\\mathbf {Y} \\mathbf {X} }=\\operatorname {K} _{\\mathbf {X} \\mathbf {Y} }^{T}} .

>>>Uncorrelatedness

Two random vectors X = ( X 1 , . . . , X m ) T {\\displaystyle \\mathbf {X} =(X_{1},...,X_{m})^{T}} and Y = ( Y 1 , . . . , Y n ) T {\\displaystyle \\mathbf {Y} =(Y_{1},...,Y_{n})^{T}} are called `!uncorrelated`! if

E ⁡ ⁡ [ X Y T ] = E ⁡ ⁡ [ X ] E ⁡ ⁡ [ Y ] T {\\displaystyle \\operatorname {E} [\\mathbf {X} \\mathbf {Y} ^{T}]=\\operatorname {E} [\\mathbf {X} ]\\operatorname {E} [\\mathbf {Y} ]^{T}} .

They are uncorrelated `F33f`_`[if and only if`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=If_and_only_if]`_`f their cross-covariance matrix K X Y {\\displaystyle \\operatorname {K} _{\\mathbf {X} \\mathbf {Y} }} is zero.`:cite-ref-gubner-3-3[`F5bf`_`[3`#cite-note-gubner-3]`_`f]

>>Correlation and cross-correlation

>>>Definitions

The `!`F33f`_`[correlation matrix`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Autocorrelation_matrix]`_`f`! (also called `!second moment`!) of an n × × 1 {\\displaystyle n\\times 1} random vector is an n × × n {\\displaystyle n\\times n} matrix whose (`*i,j`*)th element is the correlation between the `*i`* th and the `*j`* th random variables. The correlation matrix is the expected value, element by element, of the n × × n {\\displaystyle n\\times n} matrix computed as X X T {\\displaystyle \\mathbf {X} \\mathbf {X} ^{T}} , where the superscript T refers to the transpose of the indicated vector:`:cite-ref-papoulis-4-0[`F5bf`_`[4`#cite-note-papoulis-4]`_`f]`:cite-ref-gubner-3-4[`F5bf`_`[3`#cite-note-gubner-3]`_`f]

By extension, the `!cross-correlation matrix`! between two random vectors X {\\displaystyle \\mathbf {X} } and Y {\\displaystyle \\mathbf {Y} } ( X {\\displaystyle \\mathbf {X} } having n {\\displaystyle n} elements and Y {\\displaystyle \\mathbf {Y} } having p {\\displaystyle p} elements) is the n × × p {\\displaystyle n\\times p} matrix

>>>Properties

The correlation matrix is related to the covariance matrix by

R X X = K X X + E ⁡ ⁡ [ X ] E ⁡ ⁡ [ X ] T {\\displaystyle \\operatorname {R} _{\\mathbf {X} \\mathbf {X} }=\\operatorname {K} _{\\mathbf {X} \\mathbf {X} }+\\operatorname {E} [\\mathbf {X} ]\\operatorname {E} [\\mathbf {X} ]^{T}} .

Similarly for the cross-correlation matrix and the cross-covariance matrix:

R X Y = K X Y + E ⁡ ⁡ [ X ] E ⁡ ⁡ [ Y ] T {\\displaystyle \\operatorname {R} _{\\mathbf {X} \\mathbf {Y} }=\\operatorname {K} _{\\mathbf {X} \\mathbf {Y} }+\\operatorname {E} [\\mathbf {X} ]\\operatorname {E} [\\mathbf {Y} ]^{T}}

>>Orthogonality

Two random vectors of the same size X = ( X 1 , . . . , X n ) T {\\displaystyle \\mathbf {X} =(X_{1},...,X_{n})^{T}} and Y = ( Y 1 , . . . , Y n ) T {\\displaystyle \\mathbf {Y} =(Y_{1},...,Y_{n})^{T}} are called `!orthogonal`! if

E ⁡ ⁡ [ X T Y ] = 0 {\\displaystyle \\operatorname {E} [\\mathbf {X} ^{T}\\mathbf {Y} ]=0} .

>>Independence

Two random vectors X {\\displaystyle \\mathbf {X} } and Y {\\displaystyle \\mathbf {Y} } are called `!independent`! if for all x {\\displaystyle \\mathbf {x} } and y {\\displaystyle \\mathbf {y} }

F X , Y ( x , y ) = F X ( x ) ⋅ ⋅ F Y ( y ) {\\displaystyle F_{\\mathbf {X,Y} }(\\mathbf {x,y} )=F_{\\mathbf {X} }(\\mathbf {x} )\\cdot F_{\\mathbf {Y} }(\\mathbf {y} )}

where F X ( x ) {\\displaystyle F_{\\mathbf {X} }(\\mathbf {x} )} and F Y ( y ) {\\displaystyle F_{\\mathbf {Y} }(\\mathbf {y} )} denote the cumulative distribution functions of X {\\displaystyle \\mathbf {X} } and Y {\\displaystyle \\mathbf {Y} } and F X , Y ( x , y ) {\\displaystyle F_{\\mathbf {X,Y} }(\\mathbf {x,y} )} denotes their joint cumulative distribution function. Independence of X {\\displaystyle \\mathbf {X} } and Y {\\displaystyle \\mathbf {Y} } is often denoted by X ⊥ ⊥ ⊥ ⊥ Y {\\displaystyle \\mathbf {X} \\perp \\!\\!\\!\\perp \\mathbf {Y} } . Written component-wise, X {\\displaystyle \\mathbf {X} } and Y {\\displaystyle \\mathbf {Y} } are called independent if for all x 1 , … … , x m , y 1 , … … , y n {\\displaystyle x_{1},\\ldots ,x_{m},y_{1},\\ldots ,y_{n}}

F X 1 , … … , X m , Y 1 , … … , Y n ( x 1 , … … , x m , y 1 , … … , y n ) = F X 1 , … … , X m ( x 1 , … … , x m ) ⋅ ⋅ F Y 1 , … … , Y n ( y 1 , … … , y n ) {\\displaystyle F_{X_{1},\\ldots ,X_{m},Y_{1},\\ldots ,Y_{n}}(x_{1},\\ldots ,x_{m},y_{1},\\ldots ,y_{n})=F_{X_{1},\\ldots ,X_{m}}(x_{1},\\ldots ,x_{m})\\cdot F_{Y_{1},\\ldots ,Y_{n}}(y_{1},\\ldots ,y_{n})} .

>>Characteristic function

The `F33f`_`[characteristic function`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Characteristic_function_(probability_theory)]`_`f of a random vector X {\\displaystyle \\mathbf {X} } with n {\\displaystyle n} components is a function R n → → C {\\displaystyle \\mathbb {R} ^{n}\\to \\mathbb {C} } that maps every vector ω ω = ( ω ω 1 , … … , ω ω n ) T {\\displaystyle \\mathbf {\\omega } =(\\omega _{1},\\ldots ,\\omega _{n})^{T}} to a complex number. It is defined by`:cite-ref-lapidoth-2-4[`F5bf`_`[2`#cite-note-lapidoth-2]`_`f]

φ φ X ( ω ω ) = E ⁡ ⁡ [ e i ( ω ω T X ) ] = E ⁡ ⁡ [ e i ( ω ω 1 X 1 + … … + ω ω n X n ) ] {\\displaystyle \\varphi _{\\mathbf {X} }(\\mathbf {\\omega } )=\\operatorname {E} \\left[e^{i(\\mathbf {\\omega } ^{T}\\mathbf {X} )}\\right]=\\operatorname {E} \\left[e^{i(\\omega _{1}X_{1}+\\ldots +\\omega _{n}X_{n})}\\right]} .

>>Further properties

>>>Expectation of a quadratic form

One can take the expectation of a `F33f`_`[quadratic form`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Quadratic_form_(statistics)]`_`f in the random vector X {\\displaystyle \\mathbf {X} } as follows:`:cite-ref-kendrick-5-0[`F5bf`_`[5`#cite-note-kendrick-5]`_`f]

E ⁡ ⁡ [ X T A X ] = E ⁡ ⁡ [ X ] T A E ⁡ ⁡ [ X ] + tr ⁡ ⁡ ( A K X X ) , {\\displaystyle \\operatorname {E} [\\mathbf {X} ^{T}A\\mathbf {X} ]=\\operatorname {E} [\\mathbf {X} ]^{T}A\\operatorname {E} [\\mathbf {X} ]+\\operatorname {tr} (AK_{\\mathbf {X} \\mathbf {X} }),}

where K X X {\\displaystyle K_{\\mathbf {X} \\mathbf {X} }} is the covariance matrix of X {\\displaystyle \\mathbf {X} } and tr {\\displaystyle \\operatorname {tr} } refers to the `F33f`_`[trace`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Trace_(linear_algebra)]`_`f of a matrix — that is, to the sum of the elements on its `F33f`_`[main diagonal`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Main_diagonal]`_`f (from upper left to lower right). Since the quadratic form is a scalar, so is its expectation.

`!Proof`!: Let z {\\displaystyle \\mathbf {z} } be an m × × 1 {\\displaystyle m\\times 1} random vector with E ⁡ ⁡ [ z ] = μ μ {\\displaystyle \\operatorname {E} [\\mathbf {z} ]=\\mu } and Cov ⁡ ⁡ [ z ] = V {\\displaystyle \\operatorname {Cov} [\\mathbf {z} ]=V} and let A {\\displaystyle A} be an m × × m {\\displaystyle m\\times m} non-stochastic matrix.

Then based on the formula for the covariance, if we denote z T = X {\\displaystyle \\mathbf {z} ^{T}=\\mathbf {X} } and z T A T = Y {\\displaystyle \\mathbf {z} ^{T}A^{T}=\\mathbf {Y} } , we see that:

Cov ⁡ ⁡ [ X , Y ] = E ⁡ ⁡ [ X Y T ] − − E ⁡ ⁡ [ X ] E ⁡ ⁡ [ Y ] T {\\displaystyle \\operatorname {Cov} [\\mathbf {X} ,\\mathbf {Y} ]=\\operatorname {E} [\\mathbf {X} \\mathbf {Y} ^{T}]-\\operatorname {E} [\\mathbf {X} ]\\operatorname {E} [\\mathbf {Y} ]^{T}}

Hence

E ⁡ ⁡ [ X Y T ] = Cov ⁡ ⁡ [ X , Y ] + E ⁡ ⁡ [ X ] E ⁡ ⁡ [ Y ] T E ⁡ ⁡ [ z T A z ] = Cov ⁡ ⁡ [ z T , z T A T ] + E ⁡ ⁡ [ z T ] E ⁡ ⁡ [ z T A T ] T = Cov ⁡ ⁡ [ z T , z T A T ] + μ μ T ( μ μ T A T ) T = Cov ⁡ ⁡ [ z T , z T A T ] + μ μ T A μ μ , {\\displaystyle {\\begin{aligned}\\operatorname {E} [XY^{T}]&=\\operatorname {Cov} [X,Y]+\\operatorname {E} [X]\\operatorname {E} [Y]^{T}\\\\\\operatorname {E} [z^{T}Az]&=\\operatorname {Cov} [z^{T},z^{T}A^{T}]+\\operatorname {E} [z^{T}]\\operatorname {E} [z^{T}A^{T}]^{T}\\\\&=\\operatorname {Cov} [z^{T},z^{T}A^{T}]+\\mu ^{T}(\\mu ^{T}A^{T})^{T}\\\\&=\\operatorname {Cov} [z^{T},z^{T}A^{T}]+\\mu ^{T}A\\mu ,\\end{aligned}}}

which leaves us to show that

Cov ⁡ ⁡ [ z T , z T A T ] = tr ⁡ ⁡ ( A V ) . {\\displaystyle \\operatorname {Cov} [z^{T},z^{T}A^{T}]=\\operatorname {tr} (AV).}

This is true based on the fact that one can `F33f`_`[cyclically permute matrices when taking a trace`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Trace_(linear_algebra)]`_`f without changing the end result (e.g.: tr ⁡ ⁡ ( A B ) = tr ⁡ ⁡ ( B A ) {\\displaystyle \\operatorname {tr} (AB)=\\operatorname {tr} (BA)} ).

We see `F33f`_`[that`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Covariance]`_`f

Cov ⁡ ⁡ [ z T , z T A T ] = E ⁡ ⁡ [ ( z T − − E ( z T ) ) ( z T A T − − E ( z T A T ) ) T ] = E ⁡ ⁡ [ ( z T − − μ μ T ) ( z T A T − − μ μ T A T ) T ] = E ⁡ ⁡ [ ( z − − μ μ ) T ( A z − − A μ μ ) ] . {\\displaystyle {\\begin{aligned}\\operatorname {Cov} [z^{T},z^{T}A^{T}]&=\\operatorname {E} \\left[\\left(z^{T}-E(z^{T})\\right)\\left(z^{T}A^{T}-E\\left(z^{T}A^{T}\\right)\\right)^{T}\\right]\\\\&=\\operatorname {E} \\left[(z^{T}-\\mu ^{T})(z^{T}A^{T}-\\mu ^{T}A^{T})^{T}\\right]\\\\&=\\operatorname {E} \\left[(z-\\mu )^{T}(Az-A\\mu )\\right].\\end{aligned}}}

And since

( z − − μ μ ) T ( A z − − A μ μ ) {\\displaystyle \\left({z-\\mu }\\right)^{T}\\left({Az-A\\mu }\\right)}

is a `F33f`_`[scalar`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Scalar_(mathematics)]`_`f, then

( z − − μ μ ) T ( A z − − A μ μ ) = tr ⁡ ⁡ ( ( z − − μ μ ) T ( A z − − A μ μ ) ) = tr ⁡ ⁡ ( ( z − − μ μ ) T A ( z − − μ μ ) ) {\\displaystyle (z-\\mu )^{T}(Az-A\\mu )=\\operatorname {tr} \\left({(z-\\mu )^{T}(Az-A\\mu )}\\right)=\\operatorname {tr} \\left((z-\\mu )^{T}A(z-\\mu )\\right)}

trivially. Using the permutation we get:

tr ⁡ ⁡ ( ( z − − μ μ ) T A ( z − − μ μ ) ) = tr ⁡ ⁡ ( A ( z − − μ μ ) ( z − − μ μ ) T ) , {\\displaystyle \\operatorname {tr} \\left({(z-\\mu )^{T}A(z-\\mu )}\\right)=\\operatorname {tr} \\left({A(z-\\mu )(z-\\mu )^{T}}\\right),}

and by plugging this into the original formula we get:

Cov ⁡ ⁡ [ z T , z T A T ] = E [ ( z − − μ μ ) T ( A z − − A μ μ ) ] = E [ tr ⁡ ⁡ ( A ( z − − μ μ ) ( z − − μ μ ) T ) ] = tr ⁡ ⁡ ( A ⋅ ⋅ E ⁡ ⁡ ( ( z − − μ μ ) ( z − − μ μ ) T ) ) = tr ⁡ ⁡ ( A V ) . {\\displaystyle {\\begin{aligned}\\operatorname {Cov} \\left[{z^{T},z^{T}A^{T}}\\right]&=E\\left[{\\left({z-\\mu }\\right)^{T}(Az-A\\mu )}\\right]\\\\&=E\\left[\\operatorname {tr} \\left(A(z-\\mu )(z-\\mu )^{T}\\right)\\right]\\\\&=\\operatorname {tr} \\left({A\\cdot \\operatorname {E} \\left((z-\\mu )(z-\\mu )^{T}\\right)}\\right)\\\\&=\\operatorname {tr} (AV).\\end{aligned}}}

>>>Expectation of the product of two different quadratic forms

One can take the expectation of the product of two different quadratic forms in a zero-mean `F33f`_`[Gaussian`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Joint_normality]`_`f random vector X {\\displaystyle \\mathbf {X} } as follows:`:cite-ref-kendrick-5-1[`F5bf`_`[5`#cite-note-kendrick-5]`_`f]

E ⁡ ⁡ [ ( X T A X ) ( X T B X ) ] = 2 tr ⁡ ⁡ ( A K X X B K X X ) + tr ⁡ ⁡ ( A K X X ) tr ⁡ ⁡ ( B K X X ) {\\displaystyle \\operatorname {E} \\left[(\\mathbf {X} ^{T}A\\mathbf {X} )(\\mathbf {X} ^{T}B\\mathbf {X} )\\right]=2\\operatorname {tr} (AK_{\\mathbf {X} \\mathbf {X} }BK_{\\mathbf {X} \\mathbf {X} })+\\operatorname {tr} (AK_{\\mathbf {X} \\mathbf {X} })\\operatorname {tr} (BK_{\\mathbf {X} \\mathbf {X} })}

where again K X X {\\displaystyle K_{\\mathbf {X} \\mathbf {X} }} is the covariance matrix of X {\\displaystyle \\mathbf {X} } . Again, since both quadratic forms are scalars and hence their product is a scalar, the expectation of their product is also a scalar.

>>Applications

>>>Portfolio theory

In `F33f`_`[portfolio theory`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Portfolio_theory]`_`f in `F33f`_`[finance`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Finance]`_`f, an objective often is to choose a portfolio of risky assets such that the distribution of the random portfolio return has desirable properties. For example, one might want to choose the portfolio return having the lowest variance for a given expected value. Here the random vector is the vector r {\\displaystyle \\mathbf {r} } of random returns on the individual assets, and the portfolio return `*p`* (a random scalar) is the inner product of the vector of random returns with a vector `*w`* of portfolio weights — the fractions of the portfolio placed in the respective assets. Since `*p`* = `*w`*T r {\\displaystyle \\mathbf {r} } , the expected value of the portfolio return is `*w`*TE( r {\\displaystyle \\mathbf {r} } ) and the variance of the portfolio return can be shown to be `*w`*TC`*w`*, where C is the covariance matrix of r {\\displaystyle \\mathbf {r} } .

>>>Regression theory

In `F33f`_`[linear regression`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Linear_regression]`_`f theory, we have data on `*n`* observations on a dependent variable `*y`* and `*n`* observations on each of `*k`* independent variables `*xj`*. The observations on the dependent variable are stacked into a column vector `*y`*; the observations on each independent variable are also stacked into column vectors, and these latter column vectors are combined into a `F33f`_`[design matrix`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Design_matrix]`_`f `*X`* (not denoting a random vector in this context) of observations on the independent variables. Then the following regression equation is postulated as a description of the process that generated the data:

y = X β β + e , {\\displaystyle y=X\\beta +e,}

where β is a postulated fixed but unknown vector of `*k`* response coefficients, and `*e`* is an unknown random vector reflecting random influences on the dependent variable. By some chosen technique such as `F33f`_`[ordinary least squares`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Ordinary_least_squares]`_`f, a vector β β ^ ^ {\\displaystyle {\\hat {\\beta }}} is chosen as an estimate of β, and the estimate of the vector `*e`*, denoted e ^ ^ {\\displaystyle {\\hat {e}}} , is computed as

e ^ ^ = y − − X β β ^ ^ . {\\displaystyle {\\hat {e}}=y-X{\\hat {\\beta }}.}

Then the statistician must analyze the properties of β β ^ ^ {\\displaystyle {\\hat {\\beta }}} and e ^ ^ {\\displaystyle {\\hat {e}}} , which are viewed as random vectors since a randomly different selection of `*n`* cases to observe would have resulted in different values for them.

>>>Vector time series

The evolution of a `*k`*×1 random vector X {\\displaystyle \\mathbf {X} } through time can be modelled as a `F33f`_`[vector autoregression`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Vector_autoregression]`_`f (VAR) as follows:

X t = c + A 1 X t − − 1 + A 2 X t − − 2 + ⋯ ⋯ + A p X t − − p + e t , {\\displaystyle \\mathbf {X} _{t}=c+A_{1}\\mathbf {X} _{t-1}+A_{2}\\mathbf {X} _{t-2}+\\cdots +A_{p}\\mathbf {X} _{t-p}+\\mathbf {e} _{t},\\,}

where the `*i`*-periods-back vector observation X t − − i {\\displaystyle \\mathbf {X} _{t-i}} is called the `*i`*-th lag of X {\\displaystyle \\mathbf {X} } , `*c`* is a `*k`* × 1 vector of constants (`F33f`_`[intercepts`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Y-intercept]`_`f), `*Ai`* is a time-invariant `*k`* × `*k`* `F33f`_`[matrix`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Matrix_(mathematics)]`_`f and e t {\\displaystyle \\mathbf {e} _{t}} is a `*k`* × 1 random vector of `F33f`_`[error`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Errors_and_residuals_in_statistics]`_`f terms.

>>References

`:cite-note-gallager-1`!1.`! `F0af`_`[↑`#cite-ref-gallager-1-0]`_`f `:citerefgallager2013`aGallager, Robert G. (2013). `*Stochastic Processes Theory for Applications`*. Cambridge University Press. `F33f`_`[ISBN`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=ISBN_(identifier)]`_`f 978-1-107-03975-9.
`:cite-note-lapidoth-2`!2.`! `F0af`_`[↑`#cite-ref-lapidoth-2-0]`_`f `:citereftaboga2017`aTaboga, Marco (2017). `*Lectures on Probability Theory and Mathematical Statistics`*. CreateSpace Independent Publishing Platform. `F33f`_`[ISBN`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=ISBN_(identifier)]`_`f 978-1981369195.
`:cite-note-gubner-3`!3.`! `F0af`_`[↑`#cite-ref-gubner-3-0]`_`f `:citerefgubner2006`aGubner, John A. (2006). `*Probability and Random Processes for Electrical and Computer Engineers`*. Cambridge University Press. `F33f`_`[ISBN`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=ISBN_(identifier)]`_`f 978-0-521-86470-1.
`:cite-note-papoulis-4`!4.`! `F0af`_`[↑`#cite-ref-papoulis-4-0]`_`f `:citerefpapoulis1991`aPapoulis, Athanasius (1991). `*Probability, Random Variables and Stochastic Processes`* (Third ed.). McGraw-Hill. `F33f`_`[ISBN`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=ISBN_(identifier)]`_`f 0-07-048477-5.
`:cite-note-kendrick-5`!5.`! `F0af`_`[↑`#cite-ref-kendrick-5-0]`_`f `:citerefkendrick1981`aKendrick, David (1981). `*Stochastic Control for Economic Models`*. McGraw-Hill. `F33f`_`[ISBN`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=ISBN_(identifier)]`_`f 0-07-033962-7.

>>Further reading

• `:citerefstarkwoods2012`aStark, Henry; Woods, John W. (2012). "Random Vectors". `*Probability, Statistics, and Random Processes for Engineers`* (Fourth ed.). Pearson. pp. 295–339. `F33f`_`[ISBN`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=ISBN_(identifier)]`_`f 978-0-13-231123-6.

`c`F0af`_`[↑ Back to top`#top]`_`f`a